Papers with automatic speech recognition system
Sisyphus, a Workflow Manager Designed for Machine Translation and Automatic Speech Recognition (D18-2)
Copied to clipboard
| Challenge: | Sisyphus is a workflow manager for Python that can be used for large and complicated workflows. |
| Approach: | Sisyphus is a Python-based workflow manager that can be used to train and test a machine . it maps all jobs to a unique path and can create links bearing descriptive names. |
| Outcome: | Sisyphus is a Python-based workflow manager that can handle large experiments . it can be used without modification to edit, debug, document the workflow . |
TLT-school: a Corpus of Non Native Children Speech (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of speech utterances collected in schools of northern italy is being used to assess the performance of students learning both English and German. |
| Approach: | a corpus of speech utterances collected in schools of northern italy is described . the corpus is going to be freely distributed to scientific community . |
| Outcome: | The corpus of speech utterances collected in schools of northern italy is a "Trentino Language Testing" in schools" the data are used to assess the performance of students learning English and German . |
Role-specific Language Models for Processing Recorded Neuropsychological Exams (N18-2)
Copied to clipboard
| Challenge: | Neuropsychological examinations are an important screening tool for the presence of cognitive conditions such as Alzheimer's, Parkinson's and spinal-cord injuries. |
| Approach: | They propose to use audio recordings to determine the cognitive health of 92 subjects from audio that was diarized using an automatic speech recognition system trained on TED talks and on structured language used by testers and subjects. |
| Outcome: | The proposed method can determine the cognitive health of 92 subjects from audio that was diarized using an automatic speech recognition system trained on TED talks and on the structured language used by testers and subjects. |
Automatic Speech Recognition for Gascon and Languedocian Variants of Occitan (2024.lrec-main)
Copied to clipboard
Iñigo Morcillo, Igor Leturia, Ander Corral, Xabier Sarasola, Michaël Barret, Aure Séguier, Benaset Dazéas
| Challenge: | a new system for automatic speech recognition is being developed for two main Occitan dialects . the difficulty lies in the fact that Occitian is a less-resourced language . |
| Approach: | They propose to develop an automatic speech recognition system for two Occitan dialects . they use Kaldi, acoustic models, and Whisper to create a model from corpora . |
| Outcome: | The proposed system is based on Kaldi and Whisper for two main Occitan dialects . the system is more robust than previous systems, and the results are promising . |
Automatic Speech Recognition and Query By Example for Creole Languages Documentation (2022.findings-acl)
Copied to clipboard
| Challenge: | CREAM project aims to provide linguists with new methods for language documentation based on automatic speech recognition and keyword-spotting. |
| Approach: | They propose to use one hour of annotated data to design an automatic speech recognition system for two Creole languages. |
| Outcome: | The proposed model is based on an hour of annotated data and is usable by linguists. |
Does Joint Training Really Help Cascaded Speech Translation? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Currently, in speech translation, the straightforward approach delivers state-of-the-art results, but fundamental challenges such as error propagation remain. |
| Approach: | They propose to combine a cascaded recognition system with a machine translation system to improve cascade speech translation. |
| Outcome: | The proposed methods can improve cascaded speech translation and suggest alternative training methods. |
Code-Mixed Text Augmentation for Latvian ASR (2024.lrec-main)
Copied to clipboard
| Challenge: | a new study attempts to tackle code-mixed speech recognition by improving the language model of a hybrid system. |
| Approach: | They propose an inflected transliteration and phonetic transcription model for code-mixed Latvian sentences . they leverage a large human-translated English-Latvian parallel text corpus to generate synthetic Latvian phrases . |
| Outcome: | The proposed system improves on a human-translated English-Latvian parallel text corpus . the results show that the proposed system can generate code-mixed Latvian sentences . |
Data-Driven Pronunciation Modeling of Swiss German Dialectal Speech for Automatic Speech Recognition (L18-1)
Copied to clipboard
| Challenge: | a Swiss German speech recognizer is trained using a standard German annotation model. |
| Approach: | They propose to train a Swiss German speech recognition system using a standard German annotation model. |
| Outcome: | The proposed system is based on a standard German annotation model and a grapheme-to-phoneme conversion model. |
Where are we in Named Entity Recognition from Speech? (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition is usually made through a pipeline process that consists of processing audio and applying a NER to the audio outputs. |
| Approach: | They propose an original 3-pass approach and explore the capability of an E2E system to do structured NER. |
| Outcome: | The proposed system performs better than the current pipeline approach. |
Interpersonal Relationship Labels for the CALLHOME Corpus (L18-1)
Copied to clipboard
| Challenge: | a lack of corpora makes exploration of this problem intractable, says nicolaus mills . mills: communication is one of the most invaluable tools humans have . |
| Approach: | a new study uses a corpus of interpersonal relationship labels to help identify relationships . a set of labels is available for download on the website of the cnn.org team . |
| Outcome: | a new set of interpersonal relationship labels is released for the CALLHOME English corpus . the labels are available for download on the cnn.com website . |